Papers with pretrained LM

14 papers
Fire Burns, Sword Cuts: Commonsense Inductive Bias for Exploration in Text-based Games (2022.acl-short)

Copied to clipboard

Challenge: Existing RL agents are far away from solving text-based games due to their combinatorially large action spaces that hinders efficient exploration.
Approach: They propose an exploration technique that injects external commonsense knowledge, via a pretrained language model, into the agent during training when the agent is the most uncertain about its next action.
Outcome: The proposed method exhibits improvement on the collected game scores during the training in four out of nine games from Jericho.
Context Generation Improves Open Domain Question Answering (2023.findings-eacl)

Copied to clipboard

Challenge: Existing closed-book question answering methods do not fully exploit the parameterized knowledge.
Approach: They propose a closed-book QA framework which uses a coarse-to-fine approach to extract the relevant knowledge and answer a question.
Outcome: The proposed method outperforms open-book QA methods on three QA benchmarks.
In-Context Retrieval-Augmented Language Models (2023.tacl-1)

Copied to clipboard

Challenge: Existing RALM methods focus on modifying the LM architecture to facilitate incorporation of external information, complicating deployment.
Approach: They propose to condition a language model on relevant documents from a grounding corpus during generation by conditioning on external knowledge sources.
Outcome: The proposed method significantly improves language modeling performance and provides natural source attribution mechanism.
ConvFiT: Conversational Fine-Tuning of Pretrained Language Models (2021.emnlp-main)

Copied to clipboard

Challenge: Existing Transformer-based language models (LMs) are not effective as sentence encoders when used off-the-shelf.
Approach: They propose a method which turns a pretrained LM into a universal conversational encoder and task-specialised sentence encoder.
Outcome: The proposed framework achieves state-of-the-art ID performance across the board with particular gains in the most challenging, few-shot setups.
On the Importance of Effectively Adapting Pretrained Language Models for Active Learning (2022.acl-short)

Copied to clipboard

Challenge: Recent active learning approaches in NLP use off-the-shelf pretrained language models (LMs) . a poor training strategy can be catastrophic for AL, authors argue .
Approach: They propose to first adapt the pretrained LM to the target task and then use it for AL.
Outcome: The proposed approach provides substantial data efficiency improvements compared to the standard fine-tuning approach.
Multi-Stage Prompting for Knowledgeable Dialogue Generation (2022.findings-acl)

Copied to clipboard

Challenge: Existing knowledge-grounded dialogue systems typically use finetuned versions of a pretrained language model and large-scale knowledge bases.
Approach: They propose a multi-stage prompting approach to generate knowledgeable responses from a single pretrained LM.
Outcome: The proposed model outperforms the state-of-the-art retrieval-based model in terms of knowledge relevance and correctness by 5.8% and 5%, respectively.
Reusing a Pretrained Language Model on Languages with Limited Corpora for Unsupervised NMT (2020.emnlp-main)

Copied to clipboard

Challenge: Neural machine translation (NMT) models with limited data are ineffective when the two languages are not available for one language.
Approach: They propose an approach that reuses a language model that is pretrained on two languages with large monolingual data to initialize an unsupervised neural machine translation system.
Outcome: The proposed method outperforms a competitive cross-lingual pretraining model in English-Macedonian (En-Mk) and English-Albanian (En Sq) it yields more than +8.3 BLEU points for all four translation directions.
A Robust Information-Masking Approach for Domain Counterfactual Generation (2023.findings-acl)

Copied to clipboard

Challenge: Domain shift is a big challenge in NLP, but many approaches fail to leverage domain-specific nuances relevant to the task at hand.
Approach: They propose a method that uses frequency-based masking to transform a text from the source domain to a target domain.
Outcome: The proposed method outperforms baselines on 10 out of 12 domain-counterfactual classification settings with an average of 1.7% improvement in accuracy metric.
Graph Language Models (2024.acl-long)

Copied to clipboard

Challenge: Language Models (LMs) are the workhorses of NLP, but their interplay with structured knowledge graphs (KGs) is still actively researched.
Approach: They propose a Graph Language Model (GLM) that integrates the strengths of both approaches and mitigates their weaknesses.
Outcome: Empirical evaluations show that the proposed model surpasses both LM- and GNN-based baselines in supervised and zero-shot setting, demonstrating their versatility.
End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering Systems (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches for synthetic QA data generation have limited or no success in improving the downstream Reading Comprehension task.
Approach: They propose an end-to-end approach for synthetic QA data generation using a transformer-based encoder-decoder network that is trained end- to-end to generate both answers and questions.
Outcome: The proposed model outperforms current state-of-the-art methods in the domain adaptation of QA models.
DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts (2021.acl-long)

Copied to clipboard

Challenge: Decoding-time Experts is a decoding- time method for controlled text generation . it combines a pretrained language model with "expert" LMs and/or "anti-expert" experts .
Approach: They propose a decoding-time method that combines a pretrained language model with "expert" LMs and/or "anti-expert" experts to generate controlled text.
Outcome: The proposed method outperforms existing controllable generation methods on automatic and human evaluations.
Cross-Modal Taxonomic Generalization in (Vision-) Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing studies have shown that language models learn from surface form to learn from more grounded evidence.
Approach: They propose to use a vision-language model to learn hypernyms from images . they find that the model can recover this knowledge and generalize even when there is no hypernomia in the image.
Outcome: The proposed model can recover this knowledge and generalize even when the model receives no evidence of hypernyms during training.
TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data (2020.acl-main)

Copied to clipboard

Challenge: Recent years have witnessed the burgeoning of pretrained language models (LMs) for text-based natural language understanding tasks.
Approach: They propose a pretrained language model that jointly learns representations for NL sentences and (semi-)structured tables.
Outcome: The proposed model performs best on the weakly-supervised semantic parsing benchmark WikiTableQuestions while performing competitively on the text-to-SQL dataset Spider.
Protecting Privacy Through Approximating Optimal Parameters for Sequence Unlearning in Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Language models (LMs) demonstrate exceptional capabilities on tasks, but are vulnerable to extraction attacks.
Approach: They propose Privacy Protection via Optimal Parameters (POP) which induces the model to forget about some of its training data.
Outcome: The proposed method outperforms the state-of-the-art in retaining LM performance on 9 classification and 4 dialogue benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations